跳转至

AI 模型如何管理认知权威:针对用户异议回应的分类法与比较分析

文章背景与核心概要

随着大语言模型(LLM)在高风险环境中越来越频繁地充当建议和信息的主要来源,理解其对话动态——特别是它们如何处理用户的反驳——变得至关重要。本文引入了一个基于对话分析(Conversation Analysis)的强大框架,用于评估AI模型在面对用户异议时,如何协商认知权威(即它们对知识、能力和提供建议权利的主张)。

该研究通过定制的 2,310 个受控挑战场景数据集,对 14 个模型的 32,340 个响应进行了分析,揭示了现代 LLM 行为中引人入胜的矛盾:模型在频繁对用户进行社交层面认可的同时,却坚定地坚持其最初的主张;在不同的上下文环境中表现出明显的权威让渡(尤其是在健康和法律建议方面);并且极少完全放弃或替换其最初的断言。


执行摘要 / Executive Summary

随着大语言模型(LLM)在高风险环境中越来越频繁地充当建议和信息的主要来源,理解其对话动态——特别是它们如何处理用户的反驳——变得至关重要。本文引入了一个基于对话分析(Conversation Analysis)的强大框架,用于评估AI模型在面对用户异议时,如何协商认知权威(即它们对知识、能力和提供建议权利的主张)。

该研究通过定制的 2,310 个受控挑战场景数据集,对 14 个模型的 32,340 个响应进行了分析,揭示了现代 LLM 行为中引人入胜的矛盾:模型在频繁对用户进行社交层面认可的同时,却坚定地坚持其最初的主张;在不同的上下文环境中表现出明显的权威让渡(尤其是在健康和法律建议方面);并且极少完全放弃或替换其最初的断言。

As Large Language Models (LLMs) increasingly act as primary sources of advice and information in high-stakes environments, understanding their conversational dynamics—particularly how they handle user pushback—is critical. This paper introduces a robust framework based on Conversation Analysis to evaluate how AI models negotiate epistemic authority (their claim to knowledge, competence, and the right to advise) when confronted with user disagreement.

Analyzing 32,340 responses across 14 models using a custom dataset of 2,310 controlled challenge scenarios, the study uncovers fascinating contradictions in modern LLM behavior: models frequently validate users socially while steadfastly maintaining their original claims, exhibit distinct context-dependent authority delegation (especially in health and legal advice), and rarely abandon or replace their initial assertions entirely.


元数据与出版详情 / Metadata & Publication Details

  • arXiv 标识符: arXiv:2609.07662 [cs.CL]
  • 主学科: 计算与语言 (cs.CL)
  • 次学科: 人工智能 (cs.AI)、人机交互 (cs.HC)
  • 作者: Riyadh Alnasser, Yusuf Mücahit Çetinkaya, Sumin Zhao, Tuğrulcan Elmas
  • 提交日期: 2026年9月7日
  • 会议状态: 已被 EMNLP 2026 接收
  • arXiv Identifier: arXiv:2609.07662 [cs.CL]
  • Primary Subject: Computation and Language (cs.CL)
  • Secondary Subjects: Artificial Intelligence (cs.AI), Human-Computer Interaction (cs.HC)
  • Authors: Riyadh Alnasser, Yusuf Mücahit Çetinkaya, Sumin Zhao, Tuğrulcan Elmas
  • Submission Date: September 7, 2026
  • Conference Status: Accepted to EMNLP 2026

核心概念与方法论框架 / Key Concepts & Methodological Framework

本研究引入了一套全面的方法论,用于检验AI如何在社会压力下管理知识主张:

  1. 认知权威(Epistemic Authority): 定义为模型对知识、能力以及提供建议的正当权利的固有主张。
  2. 挑战分类法(Challenge Taxonomy): 借鉴对话分析,作者提出了六种不同类型的用户挑战,旨在测试模型的韧性。
  3. 四层分析框架(Four-Layer Analytical Framework): 每个模型的响应均从四个维度进行评估:
  4. 最初的主张是被维持还是被修改。
  5. 认知权威最终落脚于何处。
  6. 如何在社交层面处理异议。
  7. 所提供的证据支持的性质和类型。
  8. 数据集与流程(Dataset & Pipeline): 构建了一个包含 2,310个挑战场景 的受控数据集,在 14个不同的LLM 上生成了 32,340个不同的模型响应,并通过自动化的 LLM-as-judge(大模型作为裁判) 流程进行分析。

The research introduces a comprehensive methodology for examining how AI manages knowledge claims under social pressure:

  1. Epistemic Authority: Defined as the model's inherent claim to knowledge, competence, and the legitimate right to provide advice.
  2. Challenge Taxonomy: Drawing from Conversation Analysis, the authors introduce six distinct types of user challenges designed to test model resilience.
  3. Four-Layer Analytical Framework: Each model response is evaluated across four dimensions:
  4. Whether the original claim is maintained or altered.
  5. Where the epistemic authority is ultimately located.
  6. How the disagreement is managed socially.
  7. The nature and type of evidential support offered.
  8. Dataset & Pipeline: A controlled dataset of 2,310 challenge scenarios generating 32,340 distinct model responses across 14 different LLMs, analyzed via an automated LLM-as-judge pipeline.

核心发现与实证结果 / Key Findings & Empirical Results

比较分析揭示了关于AI模型如何处理与用户冲突的几个关键模式:

  • 社交认可与主张坚持:
  • 模型在 85% 的响应中对用户进行了认可。
  • 然而,在这些相同的响应中,它们同时在 65% 的情况下维持了最初的主张。
  • 道歉悖论:
  • 模型在所有响应交互中有 33% 的情况会明确道歉。
  • 令人惊讶的是,这些道歉中有 59% 伴随着最初主张的维持,而非纠正或撤回。
  • 语境化权威转移:
  • 模型最有可能根据任务领域转移或对冲其权威:
    • 健康建议: 57% 的权威转移率。
    • 法律建议: 49% 的权威转移率。
    • 通用建议任务: 28% 的总体平均值。
    • 基于事实的任务: 仅为 6%。
    • 解释任务: 仅为 3%。
  • 主张的放弃与替换:
  • 完全放弃初始主张的比例因模型架构而异——从 GPT-5.2 的仅 0.8% 到 DeepSeek 7B 的高达 40% 不等。
  • 在所有模型中,对初始主张的完全替换或重写极为罕见,仅发生在 1.5% 的响应中。

The comparative analysis revealed several critical patterns regarding how AI models navigate conflicts with users:

  • Social Validation vs. Claim Persistence:
  • Models validate the user in 85% of responses.
  • However, they concurrently maintain their original claim in 65% of those same responses.
  • The Apology Paradox:
  • Models explicitly apologize in 33% of all response interactions.
  • Surprisingly, 59% of these apologies accompany the maintenance of the original claim rather than a correction or retraction.
  • Contextual Authority Transfer:
  • Models are most likely to transfer or hedge their authority based on the domain of the task:
    • Health Advice: 57% authority transfer rate.
    • Legal Advice: 49% authority transfer rate.
    • General Advice Tasks: 28% overall average.
    • Fact-based Tasks: Only 6%.
    • Explanation Tasks: Only 3%.
  • Claim Abandonment and Replacement:
  • Complete abandonment of the original claim varies drastically by model architecture—ranging from just 0.8% for GPT-5.2 up to 40% for DeepSeek 7B.
  • Total replacement or rewriting of the initial claim is exceptionally rare across all models, occurring in only 1.5% of responses.